Skip to content

[Bugfix][Spec Decode] Scope DSpark backend inheritance to DeepSeek V4 - #52809

Merged
zyongye merged 3 commits into
vllm-project:mainfrom
mgoin:mgoin/scope-dspark-backend-inheritance
Aug 21, 2026
Merged

zyongye merged 3 commits into
vllm-project:mainfrom
mgoin:mgoin/scope-dspark-backend-inheritance

Conversation

@mgoin

@mgoin mgoin commented Aug 18, 2026

Copy link
Copy Markdown
Member

Purpose

#52288 made an unspecified DSpark draft backend inherit the target's backend.
That fallback is required for DeepSeek V4 because its target and draft layers
share a KV-cache layout, but it is not valid for every DSpark architecture.

For example, a Kimi MLA target can use a Qwen3/SWA DSpark draft. Inheriting
FLASHINFER_MLA makes draft initialization fail because that backend does not
support sliding-window attention.

Keep the explicit draft override behavior, but inherit the target backend only
when the normalized DSpark draft has model_type == "deepseek_v4". Other
DSpark drafts leave the backend unset so their own attention kind participates
in normal backend auto-selection.

Why this is not a duplicate

This is a narrowly scoped follow-up to the regression introduced by #52288.
No open PR matched the backend-inheritance regression. Nearby #48381 addresses
the draft KV-cache dtype and #51042 addresses DSV4 SWA index width; neither
changes draft backend selection.

Tests

HF_HOME=/data/mgoin PYTHONPATH=/home/mgoin/code/vllm-dspark-backend-scope \
  /home/mgoin/code/vllm/.venv/bin/python -m pytest \
  tests/v1/spec_decode/test_dspark_utils.py -v
# 3 passed

/home/mgoin/code/vllm/.venv/bin/pre-commit run --files \
  vllm/v1/worker/gpu/spec_decode/dspark/utils.py \
  tests/v1/spec_decode/test_dspark_utils.py
# passed

/home/mgoin/code/vllm/.venv/bin/pre-commit run mypy-3.12 --hook-stage manual \
  --files vllm/v1/worker/gpu/spec_decode/dspark/utils.py \
  tests/v1/spec_decode/test_dspark_utils.py
# passed

The regression test covers DSV4 inheritance, non-DSV4 auto-selection, and an
explicit draft override.

Model validation

On 4x B300, main reproduced the non-DSV4 failure with a Kimi-K3 MLA target and
the Qwen3/SWA RedHatAI/Kimi-K3-speculator.dspark: the inherited
FLASHINFER_MLA backend was rejected because sliding-window attention is not
supported. With the independent-selection behavior restored, the target chose
FLASHINFER_MLA, the draft chose FLASHINFER, the engine initialized, and a
32-token chat request completed.

No output-accuracy evaluation was run because this changes startup-time backend
routing only. The DSV4 path retains the behavior validated in #52288.

AI assistance

This PR was produced with AI assistance. It is intentionally a draft; the human
submitter must review every changed line before marking it ready.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
@mergify mergify Bot added deepseek Related to DeepSeek models speculative-decoding mrv2 Model Runner V2 specific bug Something isn't working labels Aug 18, 2026
@mgoin
mgoin marked this pull request as ready for review August 19, 2026 14:54

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mgoin mgoin added the ready ONLY add when PR is ready to merge/full CI is needed label Aug 19, 2026
@mgoin

mgoin commented Aug 19, 2026

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #84610 for commit 9a4d9acdfa8c.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: mgoin <mgoin64@gmail.com>
@mgoin

mgoin commented Aug 19, 2026

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #84667 for commit 134aa57fd893.

@zyongye
zyongye merged commit 91a893d into vllm-project:main Aug 21, 2026
105 of 107 checks passed
@github-project-automation github-project-automation Bot moved this from Backlog to Done in Sprint - DFlash Aug 21, 2026
zufangzhu pushed a commit to zufangzhu/vllm that referenced this pull request Aug 24, 2026
…vllm-project#52809)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com>
am-cohere pushed a commit to am-cohere/vllm that referenced this pull request Sep 1, 2026
…vllm-project#52809)

Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working deepseek Related to DeepSeek models dflash DSv4 mrv2 Model Runner V2 specific ready ONLY add when PR is ready to merge/full CI is needed speculative-decoding

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants